Papers with multimodal language

5 papers
Multimodal Language Analysis with Recurrent Multistage Fusion (D18-1)

Copied to clipboard

Challenge: Comprehending multimodal language requires modeling interactions between modalities and between them.
Approach: They propose a multistage fusion network which decomposes the fusion problem into multiple stages, each focused on a subset of multimodal signals for specialized, effective fusion.
Outcome: The proposed model performs state-of-the-art across three datasets relating to multimodal sentiment analysis, emotion recognition, and speaker traits recognition.
CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets in multimodal language are limited and disproportionately affect native speakers of other languages . authors propose a large-scale dataset for Spanish, Portuguese, German and French .
Approach: They propose a large-scale multimodal language dataset for Spanish, Portuguese, German and French.
Outcome: The proposed dataset is the largest of its kind with 40,000 total labelled sentences . it covers a diverse set topics and speakers and carries supervision of 20 labels including sentiment, emotions, and attributes.
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse? (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on the distribution of information in visually grounded contexts have focused on text-only inputs.
Approach: They propose to use multilingual vision-and-language models to estimate surprisal . they find grounding on perception increases uniformity across typologically diverse languages .
Outcome: The proposed hypothesis is tested in visual-language models over 30 languages and 13 storytelling languages . the results show grounding on perception increases uniformity across languages compared to text-only settings .
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor (D19-1)

Copied to clipboard

Challenge: Humor is a unique and creative communicative behavior often displayed during social interactions.
Approach: They present a dataset that allows to model multimodal language used in expressing humor using text, visual and acoustic communication.
Outcome: The proposed framework opens the door to understanding multimodal language used in expressing humor.
Integrating Multimodal Information in Large Pretrained Transformers (2020.acl-main)

Copied to clipboard

Challenge: Recent Transformer-based contextual word representations have shown state-of-the-art performance in multiple disciplines within NLP.
Approach: They propose an attachment to BERT and XLNet that allows them to accept multimodal nonverbal data during fine-tuning.
Outcome: The proposed attachment allows BERT and XLNet to accept multimodal nonverbal data during fine-tuning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations